Back

Computational and Structural Biotechnology Journal

American Association for the Advancement of Science (AAAS)

Preprints posted in the last 90 days, ranked by how well they match Computational and Structural Biotechnology Journal's content profile, based on 242 papers previously published here. The average preprint has a 0.23% match score for this journal, so anything above that is already an above-average fit.

1
Construction of a Standardized Time-Lapse Imaging Database and a Gradient Boosting Ensemble Framework for Integrating Zygote Morphokinetic Parameters with Conventional Embryo Assessment

ZHAO, M.; LIU, J.; HAN, D.; ZHANG, C.; ZHOU, Y.; CHEN, S.; LIU, C.

2026-08-24 obstetrics and gynecology 10.64898/2026.08.20.26359523 medRxiv
Top 0.1%
18.7%
Show abstract

In vitro fertilization (IVF) laboratories equipped with timelapse incubators generate vast quantities of sequential embryo images, yet the absence of standardized, annotated databases impedes the development of reproducible computational tools for embryo assessment. Here we describe the construction of a standardized time-lapse imaging database comprising 631 normally fertilized zygotes from 218 treatment cycles, integrating timelapse image sequences, patient clinical records, and embryo developmental outcomes. We further present a gradient boosting decision tree (GBDT) ensemble framework that integrates zygote morphokinetic parameters-continuous time-series features extracted via a validated CNN-based segmentation algorithm (US Patent US11210494B2)-with conventional embryo assessment grades (categorical features per the Istanbul consensus). The fusion framework employs equal-weight initialization followed by iterative residual-decreasing training to optimally combine heterogeneous feature types. Ablation analysis demonstrated that the integrated model achieved an AUC of 0.78, significantly outperforming morphokinetics-only (AUC 0.71) and conventional-only (AUC 0.65) models, confirming the complementary value of the two data modalities. The database and fusion framework provide a reproducible foundation for embryo development assessment and are generalizable to other multimodal data integration tasks in reproductive medicine.

2
A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides

Abhigyan, R.; Sood, V.; Arora, P.; Kaur, B.

2026-08-09 bioinformatics 10.64898/2026.08.04.742799 medRxiv
Top 0.1%
15.6%
Show abstract

Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization-based methods, which maximize prediction confidence while often deviating from the underlying distribution of experimentally validated peptides. To address this limitation, a two-phase generative-evolutionary framework is proposed that integrates distribution learning with evolutionary optimization. In the first phase, Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) are used to generate biologically plausible seed peptides. In the second phase, these peptides were used as initial seed for Hill Climbing optimization procedure that iteratively improves fitness function score. The proposed two-phase framework was evaluated using a dataset of experimentally validated IL-2-inducing peptides. Evaluation using independent IL-2 prediction models showed that Autoregressive Transformer combined with Hill Climbing achieved the best overall performance, achieving the mean IL-2 induction confidence score of 0.96 while reducing KL divergence from 2.26 for standalone Hill Climbing to 0.75. A case study on an independent IL-13 inducing peptide dataset showed similar trends, with ART initialized Hill Climbing achieving the mean IL-13 induction score of 0.99 while reducing KL divergence from 1.76 to 0.59. Overall, the framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation. HighlightsO_LIProposed a two-phase framework for bioactive peptide generation with potential to address class imbalance in peptide classification tasks. C_LIO_LIPerformed a systematic comparison of distribution-learning and optimization-based approaches for peptide generation. C_LIO_LICombined distribution-learning models for sequence generation with optimization algorithms for improving peptide functional properties. C_LIO_LIDemonstrated the applicability of the proposed framework across multiple bioactive peptide datasets. C_LI

3
A Practice on Antibody Hydrophobic Interaction Chromatography Retention Time Prediction using Pre-Trained Large Language Model Fine-Tuning

Wang, B.; Cai, B.; Chen, H.; Xia, H.; Wang, B.; Liu, J.; Han, L.; Wang, R.

2026-08-11 bioinformatics 10.64898/2026.08.05.742939 medRxiv
Top 0.1%
15.4%
Show abstract

Hydrophobicity is a critical property associated with the risk of non-specific binding, and it is commonly assessed using hydrophobic interaction chromatography retention time. Several computational approaches have been developed to predict antibody developability based on pre-trained language models. Such models can be fine-tuned with limited labeled antibody sequences and, in principle, do not require structural information, which is often challenging to obtain. Nevertheless, few studies have achieved strong performance in hydrophobicity prediction without incorporating structural features. Here, we present a case study of fine-tuning the pre-trained model IgBert to predict antibody hydrophobicity. Using Herceptin as a reference, we performed hydrophobic interaction chromatography retention time experiments and generated Herceptin-adjusted datasets. The fine-tuned model achieved a best R2 of 0.916, underscoring the critical role of rigorous data quality control. We also synthesized and validated 20 commercially available antibody sequences, and the results showed that the predicted hydrophobic properties were correctly reflected. Our findings provide practical guidance and highlight considerations for future applications of fine-tuned pre-trained language models in antibody hydrophobicity prediction. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=189 HEIGHT=200 SRC="FIGDIR/small/742939v1_ufig1.gif" ALT="Figure 1"> View larger version (36K): org.highwire.dtl.DTLVardef@9c814eorg.highwire.dtl.DTLVardef@ed609dorg.highwire.dtl.DTLVardef@62172forg.highwire.dtl.DTLVardef@1e01d37_HPS_FORMAT_FIGEXP M_FIG C_FIG

4
Bioinf-Farma: supervised integration of epitope prediction and recombinant protein developability for automated vaccine candidate prioritization

Bondi, H.; Crespi, M.; Orlando, M.; Lescai, F.; Serapian, S. A.; Colombo, G.; Fasano, M.; Pollegioni, L.; Molla, G.

2026-06-18 bioinformatics 10.64898/2026.06.15.732271 medRxiv
Top 0.1%
13.4%
Show abstract

Vaccine antigen discovery requires prioritizing protein candidates according to both immunogenic potential and recombinant expression feasibility. These properties are typically evaluated using separate computational tools, requiring researchers to integrate heterogeneous outputs through ad hoc workflows. Here, we present BIOINF-farma, a modular platform integrating epitope prediction and developability assessment for rational antigen selection within a unified environment. Candidates can be submitted as amino acid sequences or three-dimensional structures. When experimental structures are unavailable, BIOINF-farma automatically searches for models in AlphaFold DB or performs structure prediction using Boltz-2, ensuring a standardized structural representation for downstream analyses. Antigenicity is quantified by combining structure-based conformational epitope signals (MLCE/REBELOT-BEPPE) and sequence-based linear epitope propensity scores (BepiPred 3.0) into a protein-level Antigenicity Score, with a classification threshold optimized on a manually curated validation dataset. Developability is evaluated through two supervised Random Forest meta-learners that integrate three solubility predictors (DeepSoluE, SoluProt, Protein-Sol) and three thermal stability predictors (TemStaPro, ProLaTherm, BertThermo), whose outputs are combined into an Expression Efficiency Score (EES). By integrating complementary predictive signals, the meta-learning framework achieves greater accuracy and robustness than individual predictors while maintaining performance across a broad range of sequence identities. The Antigenicity Score effectively discriminates antigenic from non-antigenic proteins with a large effect size, whereas EES successfully distinguishes soluble from insoluble outcomes on an independent panel of recombinant proteins expressed in Escherichia coli. BIOINF-farma jointly assesses antigenicity and expression feasibility within a single framework. Its modular architecture facilitates the incorporation of future predictive methods, while its web-based interface makes the full pipeline accessible to users without programming expertise, supporting rapid candidate triage in vaccine research and emerging pathogen responses. Author SummaryVaccine development begins with a critical step: identifying, among the many proteins encoded in a pathogen genome, those most suitable as candidate antigens. A promising candidate must satisfy two requirements that are rarely evaluated together. It must be recognized by the immune system, so that vaccination elicits a protective response; and it must be amenable to recombinant production, since antigens that cannot be obtained in sufficient quantity and quality are of limited practical use. Current computational tools typically address only one of these aspects, and researchers must integrate their outputs manually, through procedures that are time-consuming and prone to inconsistency. We developed BIOINF-farma, an automated platform that brings these two assessments into a single analytical framework. Starting from a protein sequence or an experimental structure, the platform retrieves or predicts a three-dimensional model, evaluates the proteins antigenic potential by combining complementary epitope predictors, and estimates its expression feasibility by integrating multiple solubility and stability predictors through supervised machine learning. A web-based interface makes the full workflow available to experimental immunologists and vaccine developers without requiring computational expertise, supporting rational candidate prioritization in routine vaccine research and during emerging pathogen responses.

5
An AI System for Autonomous Algorithm Evolution in Drug Development

Zhou, Z.; Nan, Y.; Mou, M.; Qian, Y.; Liu, Y.; Zuo, Z.; Yang, H.; Xu, W.; Li, B.; Jiang, W.; Ren, Y.; Liao, Y.; Wang, Y.; Li, Y.; Yang, Q.; Xi, Z.; Mi, T.; Sun, H.; Liu, P.; Zhu, F.

2026-08-20 pharmacology and toxicology 10.64898/2026.08.16.745117 medRxiv
Top 0.1%
13.0%
Show abstract

Artificial intelligence (AI) is increasingly permeating the drug development pipeline. Numerous algorithms for accelerating this multi-stage and multi-task process have been constructed, which depends heavily on expert design and labor-intensive task-specific optimization. Given that AI-driven acceleration of drug development is recognized as a cumulative, often synergistic, effect across multiple stages, the autonomous evolution of existing algorithms across the entire pipeline is demanded to achieve a holistic advancement. Here, we present DrugEvolve, a multi-role large language model system for systematic and autonomous algorithm evolution in drug development. DrugEvolve realizes a closed-loop evolution process by incorporating Researcher, Engineer, and Analyst domains, and enables an iterative design, implementation, evaluation, and refinement of algorithm by leveraging scientific knowledge and accumulated evolutionary experience. Across eleven representative tasks spanning target identification, drug discovery, preclinical study, and clinical trial, DrugEvolve autonomously evolved the corresponding task-specific algorithms and achieved substantial performance enhancement on 120 benchmark test sets. Moreover, it showed robust generalizabilities across heterogeneous data modalities (ranging from biological sequence and graph to molecular topology and textual language), and realized gains in both predictive and generative tasks. Collectively, this AI system can serve not only as an algorithmic infrastructure for drug development, but also as a transferable paradigm for broader scientific domains.

6
Geometric characterization of the HSV - 1 glycoprotein B - amyloid β interaction in Alzheimer's disease using Forman-Ricci curvature

Bou Dagher, L.; Han, Z.; Zhou, S.; Fülöp, T.; Desroches, M.; Rodrigues, S.

2026-08-29 bioinformatics 10.64898/2026.08.26.747308 medRxiv
Top 0.1%
12.2%
Show abstract

Alzheimer's disease is characterized by the accumulation and aggregation of amyloid-{beta}(A{beta}), but the molecular mechanisms linking environmental and infectious factors to A$\beta$ conformational changes remain incompletely understood. Herpes simplex virus type 1 (HSV-1) has been proposed as a potential contributor to AD pathology, and interactions between the viral glycoprotein B (gB) and A$\beta$ may influence the conformational behaviour of the peptide. Molecular dynamics (MD) simulations provide atomic-scale information on such interactions, but conventional structural descriptors may not fully capture changes in the organization of residue interaction networks. Here, we introduce a graph-geometric framework based on Forman-Ricci curvature to characterize the evolution of residue interaction networks during MD simulations. Each simulation frame is represented as a residue interaction graph based on C--C contacts, and residue-wise curvature profiles are analysed across time. We apply the framework to A{beta}1-42 in isolation and in complex with HSV-1 gB. Conventional MD analyses indicate stable association of the simulated complex, favourable interaction energetics, and conformational changes in A{beta}, including a transition from -helical structure toward {beta}-turn-rich conformations over the simulated timescale. Forman-Ricci curvature reveals pronounced and spatially localized remodelling of the A{beta} residue interaction network in the complex, with the strongest changes concentrated in the C-terminal region. These regions also exhibit reduced temporal curvature fluctuations and progressively distinct geometric behaviour throughout the simulation. Hierarchical clustering further identifies cooperative groups of residues with coordinated curvature dynamics, including a prominent C-terminal domain. Together, these results demonstrate that Forman-Ricci curvature provides a complementary description of biomolecular dynamics by capturing changes in the geometric organization of residue interaction networks that are not directly represented by conventional structural descriptors. The framework provides a general computational approach for studying network-level structural remodelling in protein molecular dynamics and offers a quantitative perspective on the conformational consequences of HSV-1 gB--A{beta} association.

7
Sanjeevani: A manually curated anti-cancerous phytochemical database integrated with downstream analysis tools.

Jha, V.; Jha, R.; Shukla, S.; Shingan, S.; Das, G.

2026-06-19 bioinformatics 10.64898/2026.06.15.732344 medRxiv
Top 0.1%
12.1%
Show abstract

BackgroundCancer continues to pose a massive global health burden. While plant-derived phytochemicals offer promising therapeutic leads, existing natural product databases often lack cancer specificity, dataset downloadability, and integrated screening tools. MethodsWe developed Sanjeevani, an integrative web platform cataloguing 4,823 curated anticancer phytochemicals. Using a balanced dataset of 9,646 molecules, we trained Support Vector Machine (SVM), Random Forest, and K-Nearest Neighbours classifiers using a hybrid feature representation of RDKit descriptors and 2048-bit ECFP4 fingerprints. The platform also integrates AutoDock Vina for web-based molecular docking for binding affinity, poses prediction and ADMET-AI for pharmacokinetics estimation. ResultsThe SVM model demonstrated the strongest predictive capability, achieving a top test accuracy of 0.966 and a ROC-AUC of 0.992. Benchmarking across five docking tools confirmed that AutoDock Vina successfully balanced computational automation with literature-consistent binding affinity replication. The final architecture provides rapid interactive 2D/3D visualizations integrated with downstream analysis tools. ConclusionSanjeevani provides an open-access, one-stop pipeline that bridges the gap between raw natural product data and actionable computational screening, accelerating natural product-based oncology drug discovery. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/732344v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@77183borg.highwire.dtl.DTLVardef@d7e465org.highwire.dtl.DTLVardef@1d3dfd7org.highwire.dtl.DTLVardef@10cc94d_HPS_FORMAT_FIGEXP M_FIG C_FIG

8
Application of 3D Zernike Descriptors in Antibody Structural Clustering and Repurposing

de Almeida, D. d. S.; Albuquerque, A. O.; Peixoto Lima, A. M.; Gaieta, E. M.; Souza, J. S.; dos Santos-Costa, A. H.; de Andrade, L. M.; Sampaio, J. V.; Sartori, G. R.; Silva, e. J. H. M. d.

2026-08-19 bioinformatics 10.64898/2026.08.12.744489 medRxiv
Top 0.1%
12.1%
Show abstract

Antibodies generally exhibit high specificity for their cognate epitopes, but structural and physicochemical similarities between distinct epitopes can enable an antibody to recognize different antigens, resulting in cross-reactivity. This property can be exploited for antibody repurposing. To identify epitopes that share such similarities, both sequence- and structure-based approaches can be employed. In this context, 3D Zernike descriptors provide a compact representation of protein surface geometry as numerical feature vectors, enabling quantitative comparisons independently of structural alignment and orientation. Thus, this study aimed to evaluate the application of 3D Zernike descriptors for the structural clustering of antibodies and epitopes and to explore their use in antibody repurposing for the recognition of new targets. To this end, antibody binding sites previously associated with recognition of similar epitopes were analyzed at different structural levels, considering the CDRs, CDRH3, and complete paratopes. Surface similarity was subsequently quantified by calculating the Euclidean distance between their corresponding 3D Zernike feature vectors. Performance was benchmarked against SPACE2. Additionally, different distance thresholds were evaluated based on their ability to recover antibody pairs recognizing the same epitope. The paratope-based approach provided the best balance between the number of identified pairs and precision at a distance threshold of 2.7, whereas epitope clustering showed robust performance up to a distance of 3.0. At these thresholds, the 3D Zernike descriptors identified a greater number of functional pairs than SPACE2 while maintaining comparable precision and identifying complementary sets of antibody pairs.. BTaken together, these findings support the use of 3D Zernike descriptors for structural clustering of antibodies and epitopes and for guiding antibody repurposing G, a highly lethal zoonotic pathogen. Structural screening identified three antibodies with epitopes similar to the NiV target that also showed a consistent binding preference for the target epitope in molecular docking assays. Notably, one candidate, originally directed against a SARS-CoV-2 epitope, formed a stable complex with the NiV epitope, remaining within the 5 [A] RMSD threshold during heated molecular dynamics simulations and emerging as a potential cross-reactive candidate.These results support the use of this computational framework for biopharmaceutical discovery against emerging targets. Taken together, these findings support the use of 3D Zernike descriptors for structural clustering of antibodies and epitopes and for guiding antibody repurposing.

9
Cross-attention and language models reveal the interpretability of functional predictions for the human olfactory receptor family

Zhang, Y.-F.; Xu, Z.-h.; Gao, C.-x.; Duan, S.-Y.; Li, G.; Xu, C.; Lu, H.-M.

2026-08-18 bioinformatics 10.64898/2026.08.10.744067 medRxiv
Top 0.1%
12.0%
Show abstract

The attention mechanism offers the possibility for data-driven discovery of biological principles. However, for important protein families such as human olfactory receptors, the extent to which attention can associate with biologically meaningful key regions lacks systematic validation. In this study, using human olfactory receptors (ORs) as a model, we constructed CrossVOI, a VOC-OR interaction prediction framework based on protein language models and cross-attention, achieving predictive performance superior to existing methods. Furthermore, we systematically analyzed the attention distributions of CrossVOI and found that attention not only focused on ligand-binding interfaces and evolutionarily conserved sites, but also to some extent identified certain dynamically regulated regions. In summary, we propose CrossVOI, currently the best-performing framework for VOC-OR interaction prediction, and analyze the interpretability of the attention mechanism for human ORs. This study provides insights into the interpretability of protein function prediction methods and is expected to contribute to the exploration of attention mechanisms in biological mechanisms, and provide assistance for large-scale screening and mechanistic analysis of olfactory receptors.

10
Enhanced Prediction of Gut Microbiome-Related Diseases Using Hybrid Machine Learning Models

Marisetti, S. A.; Chatterjee, P.; Priyakumar, U. D.

2026-06-24 microbiology 10.64898/2026.06.24.734177 medRxiv
Top 0.1%
12.0%
Show abstract

The human gut, containing 100 trillion microbes, is also considered the "second brain," having control over the different functions of the physiological system. With advancements in bioinformatics and the development of sequencing technologies, researchers are able to explore the diversity and functional implications of gut microbiota (GM), which have become strongly associated with a variety of diseases. Microbial imbalance, or dysbiosis, acts as a biomarker for early detection and prognosis of a disease. Artificial Intelligence and Machine Learning (AI/ML) methods, although extensively used in predicting GM associated diseases, are seldom translated to having practical real-world outcomes, necessitating the design of robust AI/ML models applicable in real-world scenario. We have therefore come up with designing stacking-based ensemble architectures (EM1 and EM2), developed by integrating multiple ML-based learning algorithms for improving disease prediction accuracy. The GM datasets, after split into training and test sets, were eventually fed into the proposed two-layer ensemble models, which combines the output from standardized base learners via a meta-classifier, strengthening classification robustness as well as ensuring consistency in optimized performance across diverse datasets. Both the proposed hybrid ensemble models have emerged to be superior performers over all baseline and deep learning models, with an average accuracy of 0.87 and 0.84 respectively. By combining multiple learners, the proposed ensemble models outperform traditional single-algorithm-based approaches to attain higher accuracy and robustness on complex GM datasets. Key messagesO_LIDevelopment of stacking-based hybrid ensemble models (EM), which can be employed to integrate different AI/ML algorithms with better prediction accuracy of gut microbiome (GM)-associated diseases. C_LIO_LIUse of independent GM datasets with preprocessing methods such as SMOTE and PCA to address class imbalance and high dimensionality. C_LIO_LIAll the proposed EM architectures are mostly superior to the existing state-of-the-art AI/ML methods (highest prediction accuracy: 0.87 and 0.84 with EM1 and EM2 models respectively) for GM diseases predictions. C_LIO_LIThe cross-cohort validation demonstrates high prediction accuracy and robustness, (AUC values close to 0.98 and 0.99, for EM1 and EM2). C_LIO_LIThese therefore demonstrate the effectiveness of EM frameworks for GM associated disease prediction, paving the way for corresponding applications in precision medicine. C_LI

11
A comprehensive analysis of calreticulin mutants reveals distinct biophysicochemical proprieties with a potential for refined targeted therapies

Kurt, O. N.; Civelek, E.; Ozturk, B.; Chachoua, I.

2026-06-24 bioinformatics 10.64898/2026.06.19.733337 medRxiv
Top 0.1%
11.9%
Show abstract

Calreticulin mutations in myeloproliferative neoplasms result in the replacement of the C-terminus acidic sequence with a positively charged tail that causes pathological activation of the thrombopoietin. The two canonical variants are Type-1 and Type-2. The remaining are mainly classified as Type-1 or Type-2 like based on the wild type sequence retained. Here, we performed in silico biophysicochemical analyses of 76 CALR exon 9 frameshift variants by their sequence and predicted biophysical properties, complemented by structural modeling of the mutant homodimers. Beyond confirming the Type-1 versus Type-2 distinction, we found that the Type 1-like variants form a continuum of charge architecture along which two reproducible subgroups can be identified, rather than sharply separated classes. This work refines the conventional mechanism-based classification into a charge-resolved framework and provides testable hypotheses linking novel-tail chemistry to receptor activation in CALR-mutant neoplasms and paves the way for improved targeted therapies based on individual mutants characteristics

12
Transplanting enzyme active site geometry into antibody CDRs for catalytic antibody design

Zhu, Y.

2026-07-28 bioengineering 10.64898/2026.07.25.740676 medRxiv
Top 0.1%
11.8%
Show abstract

Antibodies provide programmable molecular recognition, whereas enzymes enable repeated chemical transformation. Catalytic antibodies seek to combine these properties within a single protein scaffold. However, conventional approaches based on transition state analogue immunisation, library screening or local mutagenesis provide limited control over the atomic arrangement of catalytic residues. They also frequently produce antibodies that bind substrates without supporting efficient chemical turnover. Recent advances in generative protein design have enabled the construction of antibody complementarity determining regions and the scaffolding of functional motifs under structural constraints. A systematic strategy for transferring experimentally supported enzyme active site geometry into antibody variable domains is still lacking. Here, we present a computational framework that treats antibody and enzyme structures as distinct but complementary inputs. Developable Fv or VHH structures provide the immunoglobulin scaffold. Enzyme complexes containing substrates, products or transition state analogues provide catalytic residues, ligand conformations, metals, cofactors and key water networks. The selected catalytic atoms are mapped into antibody complementarity determining regions, while the surrounding loops are reconstructed using antibody compatible representations and constrained all atom diffusion. Sequence design and structural back prediction are followed by filters for antibody folding, catalytic geometry, ligand positioning, conformational stability and developability. The framework avoids direct fusion of intact enzymes and antibodies. Instead, it transfers only the local geometry required for catalysis. This separation of scaffold selection from catalytic motif selection creates a testable route for determining whether natural enzyme chemistry can be embedded within antibody formats. It also provides a practical basis for evaluating substrate binding, chemical conversion, product release and catalytic turnover as separate design objectives. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=91 SRC="FIGDIR/small/740676v1_ufig1.gif" ALT="Figure 1"> View larger version (45K): org.highwire.dtl.DTLVardef@16ef6d6org.highwire.dtl.DTLVardef@f7c23org.highwire.dtl.DTLVardef@9ee50borg.highwire.dtl.DTLVardef@1cf42d4_HPS_FORMAT_FIGEXP M_FIG C_FIG

13
A Structural Antibody Benchmark of AlphaFold3 reveals Hallucinated Epitopes and a Bias for Orderness

Solanki, A.; Maurya, N. S.; Ramlakhan, M.; Li, R.; Chen, W.; Wu, Z.; Zheng, W. J.

2026-07-31 bioinformatics 10.64898/2026.07.30.741792 medRxiv
Top 0.1%
10.5%
Show abstract

AlphaFold3 has shown promise as a tool for predicting antibody-antigen binding, yet its performance across large datasets has not been fully characterized. In this study, 3401 experimentally validated antibody-antigen complexes were sourced from the Structural Antibody Database and screened alongside 23798 negative controls to benchmark AlphaFold3s binding prediction capabilities. Confidence metrics including Predicted Aligned Error and Interface Predicted Template Modeling score were used to achieving a maximum recall of 53% at 100 inference seeds. Several factors were found to influence prediction accuracy: a notable bias was observed toward antibodies derived from X-ray crystallography structures versus those from electron microscopy, and positive prediction rates were found to decrease with increasing target protein size and surface area. In contrast, neither the amino acid composition or lengths of the complementarity determining regions, nor training data leakage were found to introduce significant bias. An innate false positive rate of approximately 3% was identified, with AF3 shown to hallucinate plausible binding interfaces across the surface of decoy targets while avoiding disordered regions. Epitope mapping using DockQ, epitope shift, and antibody displacement revealed that approximately 34% of false negatives retained the correct epitope location despite poor structural alignment, suggesting that conformation refinement tools could recover additional true binding predictions. These findings provide a comprehensive characterization of AlphaFold3s strengths and limitations for antibody screening in computational drug discovery. Key MessagesO_LIAlphaFold3 has a recall of 50% and an innate false positive prediction rate of 3%. C_LIO_LIFalse negative predictions can still feature the correct epitope despite poor RMSD. C_LIO_LIFactors such as disorder and target size impact accuracy. C_LI

14
Assessing Codon Language Models for Context-Aware Codon Optimization in Nucleic Acid-Based Medicines

Toneyan, S.; Scholz, K.; De Donno, C.; Noack, F.; Auslaender, S.; Cijsouw, T.; Payne, J. L.

2026-08-19 bioinformatics 10.64898/2026.08.11.744178 medRxiv
Top 0.1%
9.9%
Show abstract

Codon optimization uses synonymous sequence changes to improve the expression and therapeutic performance of nucleic acid-based medicines. Masked language models (MLMs) have recently been proposed as alternatives to traditional, frequency-based codon optimization approaches, yet whether they offer a meaningful advantage over such simpler methods remains unclear. Here we benchmark three prominent MLMs - CaLM, EnCodon and CodonTransformer - across backtranslation fidelity, sequence generation and nine molecular phenotype prediction tasks, and experimentally evaluate model-designed sequences using a secreted embryonic alkaline phosphatase (SEAP) reporter. The models differed markedly in amino-acid fidelity and generated distinct synonymous sequence variants. However, no single model performed best across all benchmark tasks and simple sequence features remained competitive in several settings. Our interpretability analysis revealed that the models integrate a large window of codon context for making predictions, as opposed to frequency-based approaches. Our in vitro data showed that MLM-designed variants outperformed conventional and commercial-vendor-derived sequences in both transient and stably integrated expression, supporting the models ability to capture translational context beyond codon frequency. Together, our results establish MLMs as effective and complementary tools for codon optimization and suggest that sampling across multiple models may improve the likelihood of identifying high-performing therapeutic sequences.

15
Computationally mapping olfactory receptors to odor percepts using docking energy scores

Guan, Y.

2026-07-05 bioinformatics 10.64898/2026.06.30.735315 medRxiv
Top 0.1%
9.7%
Show abstract

Mapping the olfactory factors directly to perceived smells has broad implications for neuroscience, chemistry, and medicine. Since the discovery of olfactory genes in 1991, the completion of the olfactory code has been hindered by two obstacles: the unknown combinatorial principles by which ~400 receptors collectively encode thousands of perceptually distinct odors, and the absence of a complete functional map from the olfactory receptors to conscious perception. We investigated this problem by using the binding energy scores of odorants with the olfactory receptors. We first showed that using only docking scores, we can predict the smell percepts with an accuracy similar to a full set of chemical fingerprints, combining the two resulted in even better performance. This supports there is a direct relationship between olfactory receptors and specific smells. Next, we turned on the olfactory receptors one by one by iterative training and simulation, and produced corresponding specific perception profiles for each olfactory receptor. The generated matrix is sparse, with only 1-2 smell types activated for each olfactory receptors. Despite the limitation of the size of the training data, it suggests the possibility of a low-dimensional combinatorial principle underlying thousands of smells that humans can perceive. We confirmed the prediction by a list of well-known olfactory receptors. We applied the model trained on single chemicals to mixtures, confirming competitive binding was the driving force for smell specificity. This strong performance, surprisingly, is established on the simplest modeling of the binding scores on hundreds of chemicals, and we believe the mapping can become more accurate if more complicated structural modeling techniques and more data are used.

16
Systematic Benchmarking of AI-Based Molecular Generation Models for Structure-Based Drug Design

Kumar, H.; Yang, Z.; Yu, Y.; Wen, J.; Kim, P.; Zhou, X.

2026-08-20 bioinformatics 10.64898/2026.08.14.744939 medRxiv
Top 0.1%
9.6%
Show abstract

Generative artificial intelligence is accelerating molecular design, yet the relative suitability of available models for different targets and stages of preclinical drug discovery remains unclear. Here we benchmarked 12 molecular generation and optimization methods across 176 curated protein-ligand systems spanning diverse therapeutic target classes, with experimentally validated ligands providing reference chemical space. The evaluated methods encompassed pocket-conditioned 3D generation, diffusion and flow-based modeling, autoregressive construction, reference-conditioned optimization and synthesis-aware design. Performance was assessed using operational robustness, chemical validity, uniqueness, molecular and scaffold diversity, quantitative estimate of drug-likeness, synthetic accessibility, docking, physicochemical and ADMET properties, and computational resource requirements. The results revealed architecture-dependent trade off such as receptor-conditioned methods exploited binding-pocket geometry, flow-based approaches enabled efficient sampling, reference-conditioned methods favored analogue generation, and synthesis-aware approaches improved chemical feasibility, but no method consistently optimized all criteria. To address the functional potential of generated molecules, we further developed a state-aware functional classifier (SAFC) that integrates molecular dynamics derived receptor ensembles, ensemble docking and protein ligand interaction graphs. SAFC provided dynamics-aware functional activity rankings for generated molecules that were partly complementary to docking, drug-likeness and synthetic accessibility scores. These findings support hybrid, stage specific deployment of generative models rather than reliance on any single architecture or evaluation metric. This study provides practical guidelines for generative AI based preclinical drug development processes.

17
MOSurvivor-Guided Joint CpG Selection and XGBoost Hyperparameter Optimization for Compact Epigenetic Age Prediction

Yelgi, A.; Tavangari, S.; Shakarami, Z.; Janfaza, S.

2026-08-29 genomics 10.64898/2026.08.26.747213 medRxiv
Top 0.1%
9.0%
Show abstract

Accurate epigenetic age prediction from DNA methylation profiles is intrinsically high-dimensional, creating a need for parsimonious models that preserve predictive performance while reducing the number of assayed cytosine-phosphate-guanine (CpG) loci. This study introduces MOSurvivor, a population-based multi-objective search framework that jointly optimizes a weight-threshold CpG selector and eight XGBoost hyperparameters. Experiments used the GSE40279 whole-blood cohort (656 individuals profiled on the Illumina HumanMethylation450 platform). After retaining 1,000 age-correlated CpGs, five strategies were evaluated on the same 30 seeded 80:20 train/test splits: fixed-parameter XGBoost using all 1,000 CpGs, random search, a genetic algorithm, particle swarm optimization, and MOSurvivor. Internal fitness was estimated using three-fold cross-validation on each training set. Across the 30 held-out test sets, MOSurvivor achieved a mean absolute error (MAE) of 4.149 {+/-} 0.300 years, root mean squared error of 5.545 {+/-} 0.392 years, and R2 of 0.855{+/-} 0.027 while retaining 211.6 {+/-} 54.8 CpGs. Relative to full-feature XGBoost (MAE 4.095 {+/-} 0.285 years), MOSurvivor reduced the feature set by 78.8% at an MAE increase of only 0.054 years (1.3%). Paired Wilcoxon tests found no significant accuracy difference between MOSurvivor and any comparator (all unadjusted p > 0.05; all Holm-adjusted p [≥] 0.476). The most recurrent locus, cg16867657, appeared in 29 runs, whereas mean pairwise Jaccard similarity was 0.124, indicating a small stable core embedded in multiple near-equivalent feature subsets. MOSurvivor thus offers a competitive accuracy-parsimony trade-off rather than superior absolute accuracy. External validation and leakage-free nested feature preselection remain necessary before biological or clinical translation. Keywords: epigenetic clock, DNA methylation, feature selection, multi-objective optimization, XGBoost, metaheuristics, biological aging.

18
Structural bioinformatics of three Epstein-Barr Virus (EBV) Integral Membrane Proteins and their water-soluble QTY analogs

Zhang, S.; Sun, Z.; Chen, E.

2026-07-24 bioinformatics 10.64898/2026.07.22.740197 medRxiv
Top 0.1%
9.0%
Show abstract

The Epstein-Barr virus (EBV) is a highly prevalent virus worldwide that is associated with several lymphoid and epithelial malignancies. However, extensive research on EBV integral membrane proteins BILF1, LMP1 and LMP2, has been scarce due to their hydrophobic transmembrane domains. Our study applies the QTY code (glutamine, threonine, tyrosine) to design water-soluble analogs of BILF1, LMP1 and LMP2 with reduced hydrophobicity, where we systematically replaced hydrophobic amino acid residues leucine (L), isoleucine (I), valine (V), and phenylalanine (F) with structurally similar polar residues glutamine (Q), threonine (T), and tyrosine (Y). We retrieved their native sequences from UniProt, identified transmembrane domains using Protter, then performed QTY design through the Protein Solubilizing Server (PSS). We then predicted native and QTY structures using in silico prediction tools AlphaFold3, ColabFold, and Boltz-2. Our analyses demonstrate that despite significant protein sequence replacements in their transmembrane domains (54.15%-61.59%) and increased intrinsic solubility, the QTY analogs exhibited minimal changes in isoelectric point (0.00-0.15 decrease) and molecular weight (0.7-1.2 kDa increase). Additionally, structural superpositions between QTY analogs and native structures using PyMOL yield low RMSD values (0.217[A] -1.202[A]). Our results demonstrate the QTY codes ability to design detergent-free analogs of BILF1, LMP1 and LMP2 with substantially reduced hydrophobicity and aggregation propensity whilst preserving native-like structures. Our results may facilitate protein characterization studies, therapeutic research on EBV, and other protocols that typically require protein solubilization.

19
AptViralDB: A Repository of Experimentally Validated Antiviral Aptamers

Bajiya, N.; Singh, S.; Gahlot, P. S.; Raghava, G. P. S.

2026-07-11 bioinformatics 10.64898/2026.07.08.737144 medRxiv
Top 0.1%
9.0%
Show abstract

In an era of increasing drug resistance, exploring alternative molecules is crucial for the efficient management and treatment of viral diseases. Nucleic acid aptamers have emerged as highly promising candidates due to their exceptional target specificity, low immunogenicity, and versatile mechanisms for viral blocking. This manuscript describes AptViralDB, a manually curated database providing comprehensive information on experimentally validated antiviral aptamers. It contains 1,768 entries of antiviral aptamers against 40 viral species and 104 molecular targets, compiled from literature and existing databases. Each entry provides detailed annotations, including sequence, aptamer type, target, chemical modifications, binding affinity, antiviral activity, stability, and cytotoxicity. We also provide predicted secondary structures and their corresponding minimum free energy (MFE) values. Additionally, a knowledge graph created using ArcadeDB/openCypher enables users to seamlessly explore connections among aptamers, viruses, molecular targets, and biological activities. Finally, the platform offers advanced search and browsing tools, BLAST-based sequence similarity searches, GC-content analysis, downloadable datasets, and REST API access to support computational applications. (https://webs.iiitd.edu.in/raghava/aptviraldb/).

20
Real Science Is Harder Than Benchmarks: Evaluating Advanced AI Frameworks on Published Studies. I. Uncertainty Quantification, ML on Therapeutic Data Commons, and Agent-Based Modeling

Ahmed, M. O.; Amale, S. A.; Bhavsar, R. D.; Chopra, P.; Jaimes, A.; Kachhwah, A.; Kalotra, C. D.; Li, P.; Li, X.; Liao, Y.; Roy, R.; Senthilselvan, N.; Shao, Y.; Sharma, A. D.; Shrivatsan, A.; Xue, R.; You, Y.; Badkul, A.; Xie, L.; Oet, M.; Lee, K.; Sinitskiy, A.

2026-06-27 bioinformatics 10.64898/2026.06.24.734302 medRxiv
Top 0.1%
9.0%
Show abstract

Artificial Intelligence (AI) frameworks for automating scientific research have shown strong performance on benchmarks, but their capacity to routinely reproduce results from multiple real-life published studies remains largely untested. We evaluated five advanced AI research frameworks (Kosmos, K-Dense, ToolUniverse, BioAgents from bio.xyz, and the AI Scientist-v2 from Sakana AI) on three real-life tasks (including two recently published papers) spanning uncertainty quantification for molecular property predictions, machine learning on Therapeutic Data Commons benchmarks, and agent-based modeling. AI frameworks demonstrated genuine strengths: generating original hypotheses, competently executing routine data acquisition and coding tasks, providing statistical measures of confidence often absent from the original papers, and producing well-formatted final reports. At the same time, our experiments revealed that real-world scientific tasks remain considerably harder than current benchmarks suggest. No AI framework matched the scope or depth of the original studies, results varied across multiple runs of the same framework with the same prompt, and we documented cases of severe hallucinations in final reports, gaps in literature coverage, and overconfident conclusions. Verification of AI outputs required substantial domain expertise. While these three tasks are only partially representative of the broader scientific landscape, they offer a starting point for developing a more rigorous methodology for evaluation of AI performance than what is currently practiced. We conclude that AI frameworks are already valuable for prototyping research directions and stress-testing completed studies, and some of the limitations documented here appear largely tractable through infrastructure improvements and continued development.